Goto

Collaborating Authors

 manuscript id number


Joint Activity Design Heuristics for Enhancing Human-Machine Collaboration

arXiv.org Artificial Intelligence

-- Joint activity describes when more than one agent (human or machine) contributes to the completion of a task or activity. Designing for joint activity focuses on explicitly supporting the interdependencies between agents necessary for effective coordination amon g agents engaged in the joint activity. This builds and expands upon designing for usability to further address how technologies can be designed to act as effective team players. Effective joint activity requires supporting, at minimum, five primary macroc ognitive functions within teams: Event Detection, Sensemaking, Adaptability, Perspective - Shifting, and Coordination. Supporting these functions is equally as important as making technologies usable. We synthesized fourteen heuristics from relevant literatu re including display design, human factors, cognitive systems engineering, cognitive psychology, and computer science to aid the design, development, and evaluation of technologies that support joint human - machine activity . Recent advances in Artificial Intelligence (AI) and Machine Learning (ML) technologies have accelerated human - machine interactions progress ing from simple tool - based engagements to complex cognitive collaborations [1] . Machines are being designed to perform an increasing set of functions and are being expected to engage more deeply in the collaborative joint activit ies related to these functions. This shift in machine capabilities and expectations demands a corresponding re - evaluation and broadening of design and evaluation principles to support joint human - machine activity in ways that lie outside the boundaries of trad itional usability methods and models [2] . Traditional usability heuristics, such as those proposed by [3], provide a strong foundation focusing primarily on surface - level interactions such as enhancing the ease of use, efficiency, and satisfaction in human - machine interaction . These heuristics are primarily oriented towards actions and responses but offer limited support for the essential macrocognitive functions associated with effective teamwork including event detection, sensemaking, adaptability, perspective shifting, and co ordination, all of which are vital in the close collaboration of humans and machine s with joint activities [2], [4], [5], [6] . These heuristics are primarily oriented towards actions and responses but offer limited support for the essential macrocognitive functions associated with effective teamwork including event detection, sensemaking, adaptability, perspective shifting, and co ordination . A ll of these macrocognitive functions are vital in the close collaboration of humans and machines with joint activities in high - stakes and dynamic environments with little room for error [2], [5] . This reliance on macrocognitive functions is evident in domains where the ability to process complex information and adapt to changing conditions is crucial.


A Machine Learning-Driven Solution for Denoising Inertial Confinement Fusion Images

arXiv.org Artificial Intelligence

Neutron imaging is essential for diagnosing and optimizing inertial confinement fusion implosions at the National Ignition Facility. Due to the required 10-micrometer resolution, however, neutron image require image reconstruction using iterative algorithms. For low-yield sources, the images may be degraded by various types of noise. Gaussian and Poisson noise often coexist within one image, obscuring fine details and blurring the edges where the source information is encoded. Traditional denoising techniques, such as filtering and thresholding, can inadvertently alter critical features or reshape the noise statistics, potentially impacting the ultimate fidelity of the iterative image reconstruction pipeline. However, recent advances in synthetic data production and machine learning have opened new opportunities to address these challenges. In this study, we present an unsupervised autoencoder with a Cohen-Daubechies- Feauveau (CDF 97) wavelet transform in the latent space, designed to suppress for mixed Gaussian-Poisson noise while preserving essential image features. The network successfully denoises neutron imaging data. Benchmarking against both simulated and experimental NIF datasets demonstrates that our approach achieves lower reconstruction error and superior edge preservation compared to conventional filtering methods such as Block-matching and 3D filtering (BM3D). By validating the effectiveness of unsupervised learning for denoising neutron images, this study establishes a critical first step towards fully AI-driven, end-to-end reconstruction frameworks for ICF diagnostics.


Robot joint characterisation and control using a magneto-optical rotary encoder

arXiv.org Artificial Intelligence

-- A robust and compact magneto - optical rotary encoder for the characterisation of robotic rotary joints is demonstrated. The system employs magnetic field - induced optical attenuation in a double - pass configuration using rotating nonuniform magnets around an optical circulator operating in reflection . The encoder tracks continuous 360 rotation with rotation sweep rates from ฮฝ = 135 /s to ฮฝ = 3 70 /s, and an angular resolution of ฮ” ฮธ = 0. 3 . I NTRODUCTION OTARY encoders convert rotation into electromagnetic signals, most commonly electrical. Examples include precision monitoring and control of steering wheels [1], [2], motors of autopilot vehicles [2], [3], robot ics [4], [5], and prosthetic arms [6] . In robotics, the encoder is a crucial part of the positional feedback needed to perform precision movements.


STAR: A Privacy-Preserving, Energy-Efficient Edge AI Framework for Human Activity Recognition via Wi-Fi CSI in Mobile and Pervasive Computing Environments

arXiv.org Artificial Intelligence

Human Activity Recognition (HAR) via Wi - Fi Channel State Information (CSI) presents a privacy - preserving, contactless sensing approach suitable for smart homes, healthcare monitoring, and mobile IoT systems. However, existing methods often encounter comput ational inefficiency, high latency, and limited feasibility within resource - constrained, embedded mobile edge environments. This paper proposes STAR (Sensing Technology for Activity Recognition), an edge - AI - optimized framework that integrates a lightweight neural architecture, adaptive signal processing, and hardware - aware co - optimization to enable real - time, energy - efficient HAR on low - power embedded devices. STAR incorporates a streamlined Gated Recurrent Unit (GRU) - based recurrent neural netwo rk, reducing model parameters by 33% compared to conventional LSTM models while maintaining effective temporal modeling capability. A multi - stage pre - processing pipeline combining median filtering, 8th - order Butterworth low - pass filtering, and Empirical Mo de Decomposition (EMD) is employed to denoise CSI amplitude data and extract spatial - temporal features. For on - device deployment, STAR is implemented on a Rockchip RV1126 processor equipped with an embedded Neural Processing Unit (NPU), interfaced with an ESP32 - S3 - based CSI acquisition module. Experimental results demonstrate a mean recognition accuracy of 93.52% across seven activity classes and 99.11% for human presence detection, utilizing a compact 97.6k - parameter model. INT8 quantized inference achieve s a processing speed of 33 MHz with just 8% CPU utilization, delivering sixfold speed improvements over CPU - based execution. With sub - second response latency and low power consumption, the system ensures real - time, privacy - preserving HAR, offering a practi cal, scalable solution for mobile and pervasive computing environments.


maxVSTAR: Maximally Adaptive Vision-Guided CSI Sensing with Closed-Loop Edge Model Adaptation for Robust Human Activity Recognition

arXiv.org Artificial Intelligence

WiFi Channel State Information (CSI)-based human activity recognition (HAR) provides a privacy-preserving, device-free sensing solution for smart environments. However, its deployment on edge devices is severely constrained by domain shift, where recognition performance deteriorates under varying environmental and hardware conditions. This study presents maxVSTAR (maximally adaptive Vision-guided Sensing Technology for Activity Recognition), a closed-loop, vision-guided model adaptation framework that autonomously mitigates domain shift for edge-deployed CSI sensing systems. The proposed system integrates a cross-modal teacher-student architecture, where a high-accuracy YOLO-based vision model serves as a dynamic supervisory signal, delivering real-time activity labels for the CSI data stream. These labels enable autonomous, online fine-tuning of a lightweight CSI-based HAR model, termed Sensing Technology for Activity Recognition (STAR), directly at the edge. This closed-loop retraining mechanism allows STAR to continuously adapt to environmental changes without manual intervention. Extensive experiments demonstrate the effectiveness of maxVSTAR. When deployed on uncalibrated hardware, the baseline STAR model's recognition accuracy declined from 93.52% to 49.14%. Following a single vision-guided adaptation cycle, maxVSTAR restored the accuracy to 81.51%. These results confirm the system's capacity for dynamic, self-supervised model adaptation in privacy-conscious IoT environments, establishing a scalable and practical paradigm for long-term autonomous HAR using CSI sensing at the network edge.


Drone Carry-on Weight and Wind Flow Assessment via Micro-Doppler Analysis

arXiv.org Artificial Intelligence

Remote monitoring of drones has become a global objective due to emerging applications in national security and managing aerial delivery traffic. Despite their relatively small size, drones can carry significant payloads, which require monitoring, especially in cases of unauthorized transportation of dangerous goods. A drone's flight dynamics heavily depend on outdoor wind conditions and the carry-on weight, which affect the tilt angle of a drone's body and the rotation velocity of the blades. A surveillance radar can capture both effects, provided a sufficient signal-to-noise ratio for the received echoes and an adjusted postprocessing detection algorithm. Here, we conduct a systematic study to demonstrate that micro-Doppler analysis enables the disentanglement of the impacts of wind and weight on a hovering drone. The physics behind the effect is related to the flight controller, as the way the drone counteracts weight and wind differs. When the payload is balanced, it imposes an additional load symmetrically on all four rotors, causing them to rotate faster, thereby generating a blade-related micro-Doppler shift at a higher frequency. However, the impact of the wind is different. The wind attempts to displace the drone, and to counteract this, the drone tilts to the side. As a result, the forward and rear rotors rotate at different velocities to maintain the tilt angle of the drone body relative to the airflow direction. This causes the splitting in the micro-Doppler spectra. By performing a set of experiments in a controlled environment, specifically, an anechoic chamber for electromagnetic isolation and a wind tunnel for imposing deterministic wind conditions, we demonstrate that both wind and payload details can be extracted using a simple deterministic algorithm based on branching in the micro-Doppler spectra.


A Gravity-informed Spatiotemporal Transformer for Human Activity Intensity Prediction

arXiv.org Artificial Intelligence

-- Human activity intensity prediction is crucial to many location - based services. Despite tremendous p rogress in modeling d ynamics of human activity, most existing methods overlook physical constraints of spatial interaction, leading to uninterpretable spatial correlations and over - smoothing phenomenon . To address these limitations, this work proposes a physics - informed deep learning framework, namely Gravity - informed Spatiotemporal Transformer (Gravityformer) by integrat ing the universal law of gravitation to refin e transformer attention. Specifically, it (1) estimates two spatially explicit mass parameters based on spatiotemporal embedding feature, (2) models the spatial interaction in end - to - end neural network using proposed adaptive gravity model to learn the physic al constrain t, and (3) utilizes the learned spatial interaction to guide and mitigate the over - smoothing phenomenon in transformer attention. Moreover, a parallel spatiotemporal graph convolution transformer is proposed for achieving a balance between coupled spatial and temporal learning. Systematic experiments on six real - world large - scale activity datasets demonstrate the quantitative and qualitative superiority of our model over state - of - the - art benchmarks. Additionally, the learned gravity attention matrix can be not only disentangled and interpreted based on geographical laws, but also improved the generalization in zero - shot cross - region inference . This work provides a novel insight into integrating physical laws with deep learning for spatiotemporal prediction . Index Terms -- Human activity intensity prediction; Gravity model; Spatial interaction; Physics - informed machine learning; Over - smoothing phenomenon; Spatiotemporal graph neural network . This work is supported by the National Natural Science Foundation of China ( Grant # 42430106, 42371468, 424B2013) . Y i Wang, Zhenghong Wang, Fan Zhang, Chengling Tang, Weiyu Zhang and Yu Liu are with Institute of Remote Sensing and Geographic Information System, School of Earth and Space Sciences, Peking University, Beijing 100871, China. Chaogui Kang is with National Engineering Research Center of Geographic Information System, China University of Geosciences (Wuhan) 430074, China. Sijie Ruan is with School of Computer Science and Technology, Beijing Institute of Technology, Beijing 100081, China . Di Zhu and Zhongfu Ma are with Department of Geography, Environment and Society, University of Minnesota, Twin Cities, Minneapolis, MN 55455, USA . Y u Zheng is with JD iCity, JD Technology, Beijing 100176, China . P hilip S. Yu is with Department of Computer Science, University of Illinois Chicago, Chicago 60607, USA .


Human-AI Interactions: Cognitive, Behavioral, and Emotional Impacts

arXiv.org Artificial Intelligence

As stories of human-AI interactions continue to be highlighted in the news and research platforms, the challenges are becoming more pronounced, including potential risks of overreliance, cognitive offloading, social and emotional manipulation, and the nuanced degradation of human agency and judgment. This paper surveys recent research on these issues through the lens of the psychological triad: cognition, behavior, and emotion. Observations seem to suggest that while AI can substantially enhance memory, creativity, and engagement, it also introduces risks such as diminished critical thinking, skill erosion, and increased anxiety. Emotional outcomes are similarly mixed, with AI systems showing promise for support and stress reduction, but raising concerns about dependency, inappropriate attachments, and ethical oversight. This paper aims to underscore the need for responsible and context-aware AI design, highlighting gaps for longitudinal research and grounded evaluation frameworks to balance benefits with emerging human-centric risks.


A Conditional Diffusion Model for Probabilistic Prediction of Battery Capacity Degradation

arXiv.org Artificial Intelligence

Accurate prediction of lithium-ion battery capacity and its associated uncertainty is essential for reliable battery management but remains challenging due to the stochastic nature of aging. This paper presents a novel method, termed the Condition Diffusion U-Net with Attention (CDUA), which integrates feature engineering and deep learning to address this challenge. The proposed approach employs a diffusion-based generative model for time-series forecasting and incorporates attention mechanisms to enhance predictive performance. Battery capacity is first derived from real-world vehicle operation data. The most relevant features are then identified using the Pearson correlation coefficient and the XGBoost algorithm. These features are used to train the CDUA model, which comprises two core components: (1) a contextual U-Net with self-attention to capture complex temporal dependencies, and (2) a denoising network to reconstruct accurate capacity values from noisy observations. Experimental validation on the real-world vehicle data demonstrates that the proposed CDUA model achieves a relative Mean Absolute Error (MAE) of 0.94% and a relative Root Mean Square Error (RMSE) of 1.14%, with a narrow 95% confidence interval of 3.74% in relative width. These results confirm that CDUA provides both accurate capacity estimation and reliable uncertainty quantification. Comparative experiments further verify its robustness and superior performance over existing mainstream approaches.


Two-stream network-driven vision-based tactile sensor for object feature extraction and fusion perception

arXiv.org Artificial Intelligence

Tactile perception is crucial for embodied intelligent robots to recognize objects. Vision-based tactile sensors extract object physical attributes multidimensionally using high spatial resolution; however, this process generates abundant redundant information. Furthermore, single-dimensional extraction, lacking effective fusion, fails to fully characterize object attributes. These challenges hinder the improvement of recognition accuracy. To address this issue, this study introduces a two-stream network feature extraction and fusion perception strategy for vision-based tactile systems. This strategy employs a distributed approach to extract internal and external object features. It obtains depth map information through three-dimensional reconstruction while simultaneously acquiring hardness information by measuring contact force data. After extracting features with a convolutional neural network (CNN), weighted fusion is applied to create a more informative and effective feature representation. In standard tests on objects of varying shapes and hardness, the force prediction error is 0.06 N (within a 12 N range). Hardness recognition accuracy reaches 98.0%, and shape recognition accuracy reaches 93.75%. With fusion algorithms, object recognition accuracy in actual grasping scenarios exceeds 98.5%. Focused on object physical attributes perception, this method enhances the artificial tactile system ability to transition from perception to cognition, enabling its use in embodied perception applications.